Papers with evaluation protocols
Copied to clipboard
| Challenge: | This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods. |
| Approach: | This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods. |
| Outcome: | This tutorial introduces the foundations of post-training knowledge control and showcases recent frontier methods . key motivations and failure modes, harmful generation and stereotype reinforcement, are addressed . core methods such as machine unlearning, knowledge editing, and inference-time interventions are also included . |
Copied to clipboard
| Challenge: | Existing benchmarks fail to reflect robustness challenges and fairly evaluate models. |
| Approach: | They propose to ground language models to knowledge bases to investigate distribution shifts in language and linguistic aspects of distribution shift. |
| Outcome: | The proposed method fails to evaluate language models in large and small datasets . the proposed model fails to cope with unseen schemas and language variations . |
Copied to clipboard
| Challenge: | Language understanding is a multi-faceted cognitive capability, which the Natural Language Processing community has striven to model computationally for decades. |
| Approach: | They propose to rethink what constitutes tasks and model evaluation in NLP and pursue a more holistic view on language, placing trustworthiness at the center. |
| Outcome: | The proposed models are based on generative models and are being deployed in more real-world scenarios, including previously unforeseen zero-shot setups. |
Copied to clipboard
| Challenge: | Existing studies have shown that Pretrained Language Models (PLMs) perform poorly under noise due to subword segmentation. |
| Approach: | They propose a framework for subword segmentation that provides a systematic categorization of segmentation corruption under noise and evaluation protocols by generating contrastive datasets with canonical-noisy word pairs. |
| Outcome: | The proposed framework provides a systematic categorization of segmentation corruption under noise and evaluation protocols by generating contrastive datasets with canonical-noisy word pairs. |
Copied to clipboard
| Challenge: | In modern NLP, neural networks are the de-facto standard to predict complex probability measures from available context. |
| Approach: | They propose to use a single predictive distribution to evaluate models with disentangled representations of uncertainty about predictions and uncertainty about human labels. |
| Outcome: | The proposed models are crucial for trustworthy and fair NLP systems, but exploiting a single distribution is limiting. |
Copied to clipboard
| Challenge: | Existing datasets differ substantially in content distributions and annotation policies, complicating fair evaluation and generalization assessment. |
| Approach: | They quantitatively analyze dataset bias across multiple public fake news datasets with different annotation granularities, including article-level and publisher-level labels. |
| Outcome: | The proposed approach improves detection performance under in-dataset and cross-data set evaluation settings. |
Copied to clipboard
| Challenge: | Using social reasoning benchmarks, we uncover pervasive flaws in both benchmark items and evaluation methodology. |
| Approach: | They audit three widely used social reasoning benchmarks and identify flaws in their design and evaluation methodology. |
| Outcome: | The results challenge the validity of current benchmark-based claims about social reasoning in large language models. |
Copied to clipboard
| Challenge: | Existing methods for standard generation tasks fail to capture the unique dynamics of ICL. |
| Approach: | They propose a concept of self-function vectors that leverage Bayesian views and the mechanistic interpretability of ICL to model latent concept learned during in-context prompting. |
| Outcome: | The proposed framework can be used for trustworthy-related applications, such as hallucination detection. |
Copied to clipboard
| Challenge: | Existing approaches to extract relational tuples from text are incomplete and ambiguous . Existing methods rely on predefined schemas to produce t-uples . |
| Approach: | They propose a normalization-first framework that reframes OIE as a structured semantic transformation pipeline . they formalize soundness, completeness, and usefulness as approximate yet verifiable guarantees over extraction quality . |
| Outcome: | The proposed framework aims to make OIE usable for downstream reasoning and machine interpretability. |
Copied to clipboard
| Challenge: | Knowledge-based authentication is crucial for task-oriented spoken dialogue systems that offer personalised and privacy-focused services . e-learning systems should be able to enrol, identify, and verify new and recurring users based on their personal information . |
| Approach: | They propose to formalise three authentication tasks and their evaluation protocols . they propose to use a spoken multilingual dataset with 5,506 spoken dialogues . |
| Outcome: | The proposed models set the first competitive benchmarks and set directions for future research. |
Copied to clipboard
| Challenge: | Recent studies have found that large language models (LLMs) can achieve state-of-the-art performance on generic summarization benchmarks, but their performance on more complex summarizing task settings is less studied. |
| Approach: | They benchmark large language models on instruction controllable text summarization . they use 4 evaluation protocols and 11 LLMs to evaluate their performance . |
| Outcome: | The proposed model performs well on instruction controllable text summarization tasks with 4 evaluation protocols and 11 LLMs. |
Copied to clipboard
| Challenge: | Existing benchmarks and evaluation protocols suffer from inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes. |
| Approach: | They propose an automatic framework which leverages Monte Carlo Tree Search to construct numerous and diverse descriptive sentences that thoroughly represent video content in an iterative way. |
| Outcome: | The proposed framework improves MCTS-VCB and DREAM-1K on video captioning tasks by 25.0% and 16.3% respectively. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) excel at visual understanding but face severe computational bottlenecks when processing high-resolution images and long videos due to massive visual token counts. |
| Approach: | They propose a taxonomy categorizing methods into vision-side, LLM-side and hybrid paradigms and analyze token selection mechanisms and pruning strategy. |
| Outcome: | The proposed method selectively removes less informative tokens while maintaining performance. |
Copied to clipboard
| Challenge: | Existing anomaly detection methods require previous observations to be effective . contaminated observations are often not observed, making them ineffective . |
| Approach: | They propose a method that adapts a zero-shot anomaly detector to contaminated observations . they propose an evaluation suite consisting of evaluation protocols and metrics . |
| Outcome: | The proposed method adapts the zero-shot anomaly detector to contaminated observations. |
Copied to clipboard
| Challenge: | Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols . |
| Approach: | They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data . |
| Outcome: | The proposed benchmark assesses pre-trained language models on 20 diversified tasks. |
Copied to clipboard
| Challenge: | CR methods originally designed for English struggle with Morphologically Rich Languages (MRLs) a single token in Hebrew may consist of multiple anaphors, and word/morpheme boundary discrepancies make mention detection and coreference resolution difficult in MRLs. |
| Approach: | They propose a CR dataset that identifies mentions at word, sub-word and multi-word levels and an evaluation protocol that directly addresses word/morpheme boundary discrepancies. |
| Outcome: | The proposed evaluation protocol directly addresses word/morpheme boundary discrepancies in Modern Hebrew, an MRL rich with complex words and pronominal clitics. |
Copied to clipboard
| Challenge: | Recent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese–English news translation task. |
| Approach: | They empirically test neural machine translation on a Chinese–English news translation task . they show human raters prefer human over machine translation when evaluating documents . |
| Outcome: | The proposed method shows that human translators prefer document-level evaluation over machine translation . the results highlight the need to shift towards document- level evaluation as machine translation improves . |
Copied to clipboard
| Challenge: | Prior work on role-playing agents relies on supervised fine-tuning or reinforcement learning with scalarized rewards, but these approaches do not address the coordination of multiple reward dimensions during optimization. |
| Approach: | They propose a reinforcement-learning framework that enables multi-dimensional, fine-grained rubric optimization for general RPAs. |
| Outcome: | Experiments on PersonaGym and RoleMRC show that MOA improves multi-dimensional role-playing performance over supervised and standard RL baselines. |
Copied to clipboard
| Challenge: | Knowledge Base Embedding (KBE) models are widely used to encode structured information from knowledge bases, including WordNet, but the evaluation task is often focused on link prediction, ignoring their semantic capabilities. |
| Approach: | They propose to evaluate the performance of Knowledge Base Embedding (KBE) models of WordNet on link prediction and their ability to encode semantic information. |
| Outcome: | The proposed model performs poorly on two semantic tasks and two downstream tasks. |
Copied to clipboard
| Challenge: | We study whether and how cross-task generalization ability can be acquired . we use CrossFit to standardize seen/unseen task partitions and evaluation protocols . |
| Approach: | They propose a problem setup for studying cross-task generalization ability which standardizes seen/unseen task partitions and data access during different learning stages. |
| Outcome: | The proposed model can be used to build few-shot learners across diverse tasks. |
Copied to clipboard
| Challenge: | Recent detectors report near-perfect accuracy, often boasting AUROC scores above 99%, but these claims typically assume fixed generation settings, leaving open the question of how robust such systems are to changes in decoding strategies. |
| Approach: | They examine how sampling-based decoding impacts detectability with a focus on how subtle variations in a model’s (sub)word-level distribution affect detection performance. |
| Outcome: | The proposed framework systematically examines how sampling-based decoding impacts detectability, with a focus on how subtle variations in a model’s (sub)word-level distribution affect detection performance. |
Copied to clipboard
| Challenge: | Existing evaluations of large language models (LLMs) for instruction following are incomplete. |
| Approach: | They propose to use 25 base LLMs and 15 recently proposed evaluation protocols to evaluate instruction following on 4 human-annotated datasets. |
| Outcome: | The proposed evaluations identify the best-performing base LLMs and evaluation protocols with a high degree of robustness. |
Copied to clipboard
| Challenge: | Large Language Models have demonstrated remarkable capabilities in natural language understanding, reasoning, and generation. |
| Approach: | They present a comprehensive synthesis of large language models and their applications . they dissect a four-module agent architecture and review representative designs . |
| Outcome: | The proposed models address fundamental challenges in traditional recommender systems . they include limited comprehension of complex user intents, insufficient interaction capabilities . |
Copied to clipboard
| Challenge: | Existing methods for instruction tuning do not leverage the rich natural language instructions. |
| Approach: | They propose to use a benchmark to study how instruction tuning works in CL tasks. |
| Outcome: | The proposed method can achieve similar or better results than existing CL methods. |
Copied to clipboard
| Challenge: | In this paper, we reassess claims of human parity and super human performance in machine translation. |
| Approach: | They reassess claims of human parity and super human performance in machine translation . they argue that human translation involves much more than what is embedded in automatic systems . |
| Outcome: | The proposed results show that human translation involves much more than what is embedded in automatic systems. |
Copied to clipboard
| Challenge: | A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. |
| Approach: | They propose to use visual attention to build robust benchmark datasets and models that can generalize well in real-world settings. |
| Outcome: | The proposed models show that human-generated references vary drastically in different datasets/tasks, revealing the nature of each task. |
Copied to clipboard
| Challenge: | Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech. |
| Approach: | They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs . |
| Outcome: | The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks . |
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation tasks favor text generated by different LMs . human evaluation by experts is the most reliable approach, but it is costly and time-consuming . |
| Approach: | They examine whether language model-driven evaluation metrics exhibit bias toward underlying language models in the context of summarization tasks. |
| Outcome: | The proposed evaluation metrics tend to assign inflated scores to outputs generated by the very model they are based on. |
Copied to clipboard
| Challenge: | Existing studies on understanding and reasoning with abstractive information from the visual modality have not explored the use of STructured and Abstractive Reasoning (STAR) on such data. |
| Approach: | They propose an automatic STAR data engine to synthesize images with MMRK to build multi-modal instructions with reliable chain-of-thought thinking for various STAR tasks. |
| Outcome: | The proposed framework outperforms GPT-4o in STAR and improves performance across 8 open-source MLLMs. |
Copied to clipboard
| Challenge: | Abstraction and Reasoning Corpus and ARC-AGI are widely used to assess progress in artificial intelligence. |
| Approach: | They propose a two-stage pipeline that separates perception and reasoning . they propose to test this pipeline against standard end-to-end one-stage evaluation . |
| Outcome: | The proposed pipeline separates perception and reasoning, and isolates reasoning from bottlenecks. |
Copied to clipboard
| Challenge: | Existing literature on large language models (LLMs) define knowledge as a fact if it correctly completes a cloze sentence . but the predictions of semantically equivalent clozing sentences are inconsistent . |
| Approach: | They review standard definitions of knowledge in epistemology and formalize interpretations applicable to LLMs. |
| Outcome: | The authors compare the preferences of philosophers and computer scientists in terms of knowledge definitions and evaluation protocols for testing knowledge in accordance with the most relevant definitions. |
Copied to clipboard
| Challenge: | Existing evaluation protocols and metrics do not capture the full spectrum of LLM capabilities, especially in complex reasoning tasks. |
| Approach: | They propose a new evaluation metric that continuously assesses model performance across multiple sampling attempts, quantifying both the model’s potential capabilities and operational consistency. |
| Outcome: | The proposed evaluation metric measures model performance across multiple sampling attempts and provides comprehensive insights into their potential capabilities and operational consistency. |
Copied to clipboard
| Challenge: | despite advances in CRSs, reliably assessing their ability to elicit preferences remains a challenge. |
| Approach: | They propose a user-CRS evaluation protocol with target-free user simulators . they show that current evaluation metrics emphasize single-turn recall of target items . |
| Outcome: | The proposed evaluation protocol is based on a simulation-based evaluation environment. |
Copied to clipboard
| Challenge: | Understanding and controlling behavior of large language models (LLMs) is an important topic in multilingual NLP. |
| Approach: | They propose a lightweight parallel-question benchmark for evaluating language-forcing behavior in large language models across 32 languages. |
| Outcome: | The proposed benchmark measures language steering in 32 languages across 32 languages. |
Copied to clipboard
| Challenge: | In this paper, we introduce a new far-field speaker recognition benchmark called RoboVox. |
| Approach: | They introduce a new far-field speaker recognition benchmark called RoboVox which measures the far-feet of a French corpus recorded by a mobile robot. |
| Outcome: | The proposed benchmarks show a significant decline in far-field speaker recognition and urge the community to further research in this domain. |
Copied to clipboard
| Challenge: | Recent work has leveraged probing-based approaches to study the separability of malicious and benign inputs in Large Language Models’ internal representations. |
| Approach: | They propose to use probing-based methods to study separability of malicious and benign inputs in LLMs' internal representations to detect harmful and benign content. |
| Outcome: | The proposed methods show that they learn superficial patterns rather than semantic harmfulness. |
Copied to clipboard
| Challenge: | Existing text-rich image understanding benchmarks lack scale and fragmented scenarios . a new full-image structured output format is proposed to enable fine-grained evaluation of perception and reasoning capabilities. |
| Approach: | They propose a large-scale, multilingual benchmark that includes over 100,000 annotations and 22,000 question-answer pairs. |
| Outcome: | The proposed framework provides a comprehensive platform for developing and evaluating next-generation multimodal AI systems. |
Copied to clipboard
| Challenge: | Existing E2ESD benchmarks are limited by coarse-grained requirement specifications and unreliable evaluation protocols. |
| Approach: | They propose a benchmark to assess whether generated software meets user needs . they use a fine-grained set of user requirements and a fully automated testing pipeline . |
| Outcome: | E2EDev is a benchmark to assess whether generated software meets user needs through mimicking real user interactions. |
Copied to clipboard
| Challenge: | Existing approaches lack robustness to handle complex edge cases and generalizability across different domains. |
| Approach: | They develop an accurate and lightweight verifier model for evaluation and outcome reward that matches unstructured outputs against standard answers. |
| Outcome: | The proposed model can process multiple answer types including multi-subproblems, formulas, and sequence answers while identifying abnormal/invalid responses. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly adopted as scalable judges for open-ended generation, yet how they form judgments remains insufficiently understood. |
| Approach: | They show that exposing reasoning influences LLM-based judgment . they also show that reasoning fluency and factuality critically shape judgment outcomes . |
| Outcome: | Empirical results show that the presence of reasoning significantly alters judgment behavior . stronger judges exhibit more selective behavior and achieve higher judgment accuracy . |
Copied to clipboard
| Challenge: | Existing multimodal reasoning benchmarks for large vision-language models emphasize single-image analysis and fail to exploit contextual information across multiple images. |
| Approach: | They propose a benchmark to evaluate Olympiad-level reasoning when evidence is distributed over multiple images. |
| Outcome: | The proposed model outperforms existing models on bi-image Olympiads and Gemini-3-Pro on multimodal Olympiad-level reasoning tasks. |